Skip to content

[Feature][MRV2] Adapt PCP for data parallelism - #15969

Merged
weijinqian0 merged 4 commits into
vllm-project:mainfrom
wzx0726:MRV2-PCP-DP
Sep 9, 2026
Merged

weijinqian0 merged 4 commits into
vllm-project:mainfrom
wzx0726:MRV2-PCP-DP

Conversation

@wzx0726

@wzx0726 wzx0726 commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP partitioning. Saved PCP attention state may be absent or belong to the preceding real batch. Runtime dummy inputs use runner buffers, while decode graphs capture persistent PCP buffers.

  • Build dummy PCP attention contexts from the current batch, block tables and rank-local slot mappings, including the DSA metadata path.
  • Refresh persistent PCP buffers before dummy execution. Keep the implementation in AscendPCPManager; NPUModelRunner.prepare_dummy_attn delegates or takes the existing non-PCP path.
  • Synchronize replicated speculative drafts independently of the target's PCP-local DP state, with every DP replica participating, including idle and decode replicas.
  • Restrict PCP+DP graph configuration to eager or FULL_DECODE_ONLY.

Upstream validation dependency: vllm-project/vllm#54523, merged as 7c2f1ff4958eaf0818405e9192c71608fe4a16b1, moves the PCP+DP restriction from common ParallelConfig validation to CUDA/ROCm platform checks. This PR relies on that change and contains no Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867 ([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP prepare_inputs_to_capture entry point to create AscendInputBatch directly in persistent PCP buffers. After capture and real-to-idle replay validation, remove AscendPCPManager.prepare_dummy_attn and the runner override, as tracked by TODO(wzx0726). The Ascend dummy attention-context handling and replicated-draft DP synchronization remain necessary.

Does this PR introduce any user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI options or environment variables are introduced.

Rebased onto Ascend main ad86348b0cb2324d643df6a496f2d9c3879e4481. The draft conflict resolution preserves vLLM 0.28.0 num_tokens_across_dp forwarding and applies fresh synchronization for replicated drafts on the main2main dp_sync interface.

Compatibility and remaining validation:

  • The previously tested baseline was Ascend b1c91857c16bde10c2fa7f5d7548b7666a86d4bb with vLLM e6bfe03ad73a3330cb427885aa90d97a12e1c704.
  • Removing the validator bypass requires vLLM #54523. Ascend main currently pins b2f685834a6456197e7033966fdef52a23f1abcd, which predates that change and still rejects PCP+DP. The local vLLM checkout remains at the user-selected 7c2f1ff4958eaf0818405e9192c71608fe4a16b1. This PR does not change the main2main pin or reintroduce the bypass.
  • Compatibility work against that newer revision is still pending, including the added prepare_dummy_attn argument and removal of InputBatch.max_seq_len_np. The current branch is not yet validated to run against that revision.
  • Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions during FULL_DECODE_ONLY replay, speculative-decoding numerical correctness, and EP/TP combinations require model-level validation on the final paired sources.

How was this patch tested?

The latest rebase onto ad86348b0 resolves an append-location conflict in test_pcp_manager_v2.py, retaining both the main DCP padding regression and this PR's tests. AST comparison confirms that all main tests and the final PR tests are preserved, and the other changed files match the automatic merge. Ruff lint/format, syntax parsing, and git diff --check passed. Runtime tests were not rerun for this conflict-only update.

For the preceding rebase onto f8481287a, 6 targeted candidate-method cases passed in an isolated process in the Ascend container: five main2main draft-sync cases and one mocked v0.28.0 argument-routing case. The exact candidate propose method was loaded into the installed runtime classes; native dependencies and the rest of the runtime remained at their installed versions. This is method-level regression evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax parsing, and git diff --check. Full unit-suite and model validation on the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline in an isolated Ascend-container test directory, using CPU tensors and mocks for device kernels:

python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
  • 75 passed (14 existing torch.jit deprecation warnings). This is prior-baseline evidence, not a test result for vLLM 7c2f1ff.

  • Covers missing/stale dummy PCP state, rank-local mappings, persistent buffer contents and storage, DSA metadata forwarding, non-PCP fallback, graph restrictions, and replicated draft DP synchronization.

  • The removed validator bypass's dedicated tests are also removed; the existing unrelated platform test is retained.

  • For the removal, targeted Ruff lint/format checks, Python syntax checks and git diff --check passed. No runtime tests were rerun against the upgraded vLLM. The local format.sh ci entry point remains unavailable because its shell lacks pre-commit.

  • vLLM main: vllm-project/vllm@b2f6858

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request adapts the Ascend MRV2 PCP (Prefill Context Parallelism) to support data parallelism. It addresses discrepancies between DP metadata and partitioned model inputs by calculating padded execution lengths prior to dispatch and ensuring that input preparation correctly restores token totals. Additionally, it improves robustness by enabling the construction of dummy attention contexts for runtime batches that bypass standard partitioning.

Highlights

  • PCP and Data Parallelism Integration: Implemented logic to compute PCP-padded execution lengths before DP dispatch, ensuring DP metadata remains consistent with partitioned model inputs.
  • Dummy Attention Context Construction: Added support for building runtime dummy attention contexts from current dummy batches, block tables, and slot mappings to handle cases where saved PCP state is missing or stale.
  • Configuration Constraints: Restricted PCP+DP graph configurations to eager mode or FULL_DECODE_ONLY to ensure stability.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements support for Prefill Context Parallelism (PCP) combined with Data Parallelism (DP) in Model Runner V2, including dispatch size calculations, config validation, and dummy batch handling for attention context. Feedback highlights two critical issues: first, a potential race condition/stale data issue in gather_batch_req_state where num_computed_tokens_np is read before synchronization when speculative decoding is active; second, a potential RuntimeError in build_attention_context due to calling .view() on a potentially non-contiguous slot_mappings tensor, which should be replaced with .reshape().\n\nSuggested PR Title:\n\nmarkdown\n[Ops][Feature] Support PCP+DP in Model Runner V2 and handle dummy batches\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis PR introduces support for Prefill Context Parallelism (PCP) combined with Data Parallelism (DP) in Model Runner V2. It adds a mechanism to query PCP execution length for dispatch, updates input preparation to handle PCP-local dispatch totals during prefill, and enhances `build_attention_context` to support runtime dummy batches. Additionally, it restricts PCP+DP to eager mode or FULL_DECODE_ONLY CUDA graphs.\n\n### Does this PR introduce _any_ user-facing change?\nNo.\n\n### How was this patch tested?\nTested via new unit tests in `test_attn_utils_v2.py`, `test_model_runner_v2.py`, and `test_pcp_manager_v2.py`.\n

Comment thread vllm_ascend/worker/v2/model_runner.py Outdated
Comment on lines +568 to +582
if (
self.pcp_manager is not None
and self.dp_size > 1
and batch_state is not None
and batch_state.has_prefill
):
assert isinstance(self.pcp_manager, AscendPCPManager)
# Dispatch needs the PCP-padded execution total; per-request
# counts remain global and prepare_inputs restores their sum.
num_tokens = self.pcp_manager.get_num_tokens_for_dispatch(
batch_state.num_scheduled_tokens,
self.req_states.num_computed_tokens_np[batch_state.idx_mapping_np],
batch_state.is_prefilling_np,
)
batch_state = batch_state._replace(num_tokens=num_tokens)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

When speculative decoding is enabled (self.speculator is not None), the GPU-side num_computed_tokens may have been reverted/corrected due to token rejections. This corrected state is copied asynchronously to self.num_computed_tokens_cpu at the end of the previous step, but it is not synchronized and copied to self.req_states.num_computed_tokens_cpu (which updates num_computed_tokens_np) until _update_seq_lens_cpu is called inside prepare_inputs. Since gather_batch_req_state is called before prepare_inputs, reading self.req_states.num_computed_tokens_np here will use stale computed token counts, leading to an incorrect PCP dispatch size calculation. We should synchronize and update the CPU computed tokens for the current batch's requests before calling get_num_tokens_for_dispatch.

            if (
                self.pcp_manager is not None
                and self.dp_size > 1
                and batch_state is not None
                and batch_state.has_prefill
            ):
                assert isinstance(self.pcp_manager, AscendPCPManager)
                if self.speculator is not None:
                    self.num_computed_tokens_event.synchronize()
                    for req_id in batch_state.req_ids:
                        req_index = self.req_states.req_id_to_index[req_id]
                        self.req_states.num_computed_tokens_cpu[req_index] = self.num_computed_tokens_cpu[req_index]
                # Dispatch needs the PCP-padded execution total; per-request
                # counts remain global and prepare_inputs restores their sum.
                num_tokens = self.pcp_manager.get_num_tokens_for_dispatch(
                    batch_state.num_scheduled_tokens,
                    self.req_states.num_computed_tokens_np[batch_state.idx_mapping_np],
                    batch_state.is_prefilling_np,
                )
                batch_state = batch_state._replace(num_tokens=num_tokens)

Comment on lines +352 to +354
global_slot_mappings=slot_mappings.view(slot_mappings.shape[0], self.pcp_world_size, num_tokens)[
:, self.pcp_rank
],

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Using .view() on slot_mappings can raise a RuntimeError if the tensor is non-contiguous (e.g., if it was created via slicing or other non-contiguous operations). Since slot_mappings is passed from external callers and its contiguity is not guaranteed, it is safer to use .reshape() instead of .view().

Suggested change
global_slot_mappings=slot_mappings.view(slot_mappings.shape[0], self.pcp_world_size, num_tokens)[
:, self.pcp_rank
],
global_slot_mappings=slot_mappings.reshape(slot_mappings.shape[0], self.pcp_world_size, num_tokens)[
:, self.pcp_rank
],

@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@wzx0726
wzx0726 marked this pull request as ready for review September 8, 2026 07:27
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@Ronald1995 Ronald1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@Ronald1995 Ronald1995 added the ready-precise run selected e2e test for pr label Sep 9, 2026
Build PCP dummy attention context from the current batch and refresh
persistent graph buffers in AscendPCPManager. Keep the runner override
as a thin delegation and synchronize replicated drafts independently
of the target PCP-local DP state. Restrict graphs to FULL_DECODE_ONLY.

Add regression coverage for idle state, DSA metadata, persistent buffers,
graph modes and draft synchronization. 100 focused unit tests pass;
multi-card inference and graph replay remain pending in the draft PR.

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Remove duplicate upstream dispatch and unchanged input-preparation tests.
Reduce graph and draft parameter matrices while preserving idle context,
persistent buffer refresh, non-PCP fallback and draft DP synchronization.

Validation: 75 focused tests passed in an isolated Ascend container.
Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Preserve v0.28.0 token-count forwarding while retaining main2main replicated draft synchronization coverage. Skip DPSyncState-specific tests on the release interface. Six candidate-method cases passed in an isolated Ascend-container process with installed runtime dependencies; this is not full paired-source validation.

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
@weijinqian0
weijinqian0 merged commit 198b694 into vllm-project:main Sep 9, 2026
22 checks passed
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
weiguihua2 pushed a commit that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [#14068][ascmoon],
[#14699][ascextract] and [#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[#16306](#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend #14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend #14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend #14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend #14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend #14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend #15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend #15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend #14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Signed-off-by: tianming2009 <13246728590@163.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Signed-off-by: like-0517 <ithwlike@126.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
weijinqian0 pushed a commit that referenced this pull request Sep 30, 2026
…dling (#17632)

### What this PR does / why we need it?

The paired vLLM (`ced6857afa0ea7b2e3f0846a62e1394e90f15607`, v0.30.0)
now supports speculative PCP partitioning and stages dummy inputs into
persistent PCP buffers. Remove the compatibility paths introduced for
these upstream gaps:

- Delete `_partition_speculative_batch_compat` and call upstream
`partition_batch` directly. For decode-only batches, preserve the
attention state computed before partitioning, after refreshing local
sequence lengths and graph padding. Upstream clears local draft
metadata; recomputing the state from that metadata can turn an MTP
verification batch into `PrefillCacheHit` when chunked prefill is
disabled.
- Delete `AscendPCPManager.prepare_dummy_attn` and the
`NPUModelRunner.prepare_dummy_attn` override. Inherit the upstream
runner implementation and retain the existing Ascend PCP block-table and
attention-context extensions. Upstream `prepare_inputs_to_capture`
already refreshes the persistent device input buffers.

This follows up on #14960 and #15969. Only two production files change;
the separate E2E changes in #17150 are excluded. No new helpers, state,
flags, or parallel implementations are introduced.

### Does this PR introduce _any_ user-facing change?

No new API or configuration. PCP + MTP continues to use the existing
speculative attention state while relying on the paired upstream
implementation.

### How was this patch tested?

- Static checks: Python AST parsing and `git diff --check` passed. The
isolated PR patch matches the reviewed production changes exactly.
- The speculative-partition/state-preservation change was previously
compared against the original compatibility helper on Ascend A5: MRV2,
TP=1, PCP=2, DCP=1, MTP=3, eager, block size 128, a five-layer DeepSeek
model, chunked prefill both enabled and disabled. Single-request and
two-request batches were repeated twice with temperature 0 and seed 7;
all 768 compared output tokens matched exactly. The previously failing
non-chunked case retained `SpecDecoding`.
- The subsequent dummy-attention cleanup has **not** completed runtime
validation. The DP-idle/FULL_DECODE_ONLY comparison stopped at an
environment import failure before inference; further runs were paused at
the author's request. The earlier eager results do not validate this
cleanup or the final combined patch.
- No test files were changed. DP real/idle/real transitions,
FULL_DECODE_ONLY replay, mixed prefill/decode traffic, and full-model
coverage remain pending. This PR is intentionally a draft.

- vLLM main:
vllm-project/vllm@ced6857

Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
HackClawMxw pushed a commit to HackClawMxw/vllm-ascend-fork that referenced this pull request Sep 30, 2026
…dling (vllm-project#17632)

### What this PR does / why we need it?

The paired vLLM (`ced6857afa0ea7b2e3f0846a62e1394e90f15607`, v0.30.0)
now supports speculative PCP partitioning and stages dummy inputs into
persistent PCP buffers. Remove the compatibility paths introduced for
these upstream gaps:

- Delete `_partition_speculative_batch_compat` and call upstream
`partition_batch` directly. For decode-only batches, preserve the
attention state computed before partitioning, after refreshing local
sequence lengths and graph padding. Upstream clears local draft
metadata; recomputing the state from that metadata can turn an MTP
verification batch into `PrefillCacheHit` when chunked prefill is
disabled.
- Delete `AscendPCPManager.prepare_dummy_attn` and the
`NPUModelRunner.prepare_dummy_attn` override. Inherit the upstream
runner implementation and retain the existing Ascend PCP block-table and
attention-context extensions. Upstream `prepare_inputs_to_capture`
already refreshes the persistent device input buffers.

This follows up on vllm-project#14960 and vllm-project#15969. Only two production files change;
the separate E2E changes in vllm-project#17150 are excluded. No new helpers, state,
flags, or parallel implementations are introduced.

### Does this PR introduce _any_ user-facing change?

No new API or configuration. PCP + MTP continues to use the existing
speculative attention state while relying on the paired upstream
implementation.

### How was this patch tested?

- Static checks: Python AST parsing and `git diff --check` passed. The
isolated PR patch matches the reviewed production changes exactly.
- The speculative-partition/state-preservation change was previously
compared against the original compatibility helper on Ascend A5: MRV2,
TP=1, PCP=2, DCP=1, MTP=3, eager, block size 128, a five-layer DeepSeek
model, chunked prefill both enabled and disabled. Single-request and
two-request batches were repeated twice with temperature 0 and seed 7;
all 768 compared output tokens matched exactly. The previously failing
non-chunked case retained `SpecDecoding`.
- The subsequent dummy-attention cleanup has **not** completed runtime
validation. The DP-idle/FULL_DECODE_ONLY comparison stopped at an
environment import failure before inference; further runs were paused at
the author's request. The earlier eager results do not validate this
cleanup or the final combined patch.
- No test files were changed. DP real/idle/real transitions,
FULL_DECODE_ONLY replay, mixed prefill/decode traffic, and full-model
coverage remain pending. This PR is intentionally a draft.

- vLLM main:
vllm-project/vllm@ced6857

Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants